Papers with cultural grounding
FFE-Hallu: Hallucinations in Fixed Figurative Expressions: A Benchmark of Idioms and Proverbs in the Persian Language (2026.eacl-long)
Copied to clipboard
| Challenge: | Figurative language, especially fixed figurative expressions, poses unique challenges for large language models . Unlike literal phrases, FFEs are culturally grounded and often non-compositional, making them vulnerable to figurativ hallucination . |
| Approach: | They propose a benchmark to evaluate LLMs' ability to generate, detect, and translate fixed figurative expressions in Persian. |
| Outcome: | The proposed benchmarks show that LLMs still struggle with figurative language expressions . the benchmarks are based on 600 carefully curated examples spanning three tasks . |
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)
Copied to clipboard
Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, Baraah Qawasmeh, Houdaifa Atou, Serry Sibaee, Hamzah A. Alsayadi, Walid Al-Dhabyani, Maged S. Al-shaibani, Aya El aatar, Nour Qandos, Rahaf Alhamouri, Samar Ahmad, Mohammed Anwar AL-Ghrawi, Aminetou Yacoub, Ruwa AbuHweidi, Vatimetou Mohamed Lemin, Reem Abdel-Salam, Ahlam Bashiti, Adel Ammar, Aisha Alansari, Ahmed Ashraf, Nora Alturayeif, Alcides Alcoba Inciarte, AbdelRahim A. Elmadany, Mohamedou Cheikh Tourad, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | a new task to evaluate text-to-image generation models for multicultural scenes is unexplored. |
| Approach: | They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior . |
| Outcome: | The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness. |